Barcoding-free BAC Pooling Enables Combinatorial Selective Sequencing of the Barley Gene Space

نویسندگان

  • Stefano Lonardi
  • Denisa Duma
  • Matthew Alpert
  • Francesca Cordero
  • Marco Beccuti
  • Prasanna Bhat
  • Yonghui Wu
  • Gianfranco Ciardo
  • Burair Alsaihati
  • Yaqin Ma
  • Steve Wanamaker
  • Josh Resnik
  • Timothy J. Close
چکیده

We propose a new sequencing protocol that combines recent advances in combinatorial pooling design and second-generation sequencing technology to efficiently approach de novo selective genome sequencing. We show that combinatorial pooling is a cost-effective and practical alternative to exhaustive DNA barcoding when dealing with hundreds or thousands of DNA samples, such as genome-tiling generich BAC clones. The novelty of the protocol hinges on the computational ability to efficiently compare hundreds of million of short reads and assign them to the correct BAC clones so that the assembly can be carried out clone-by-clone. Experimental results on simulated data for the rice genome show that the deconvolution is extremely accurate (99.57% of the deconvoluted reads are assigned to the correct BAC), and the resulting BAC assemblies have very high quality (BACs are covered by contigs over about 77% of their length, on average). Experimental results on real data for a gene-rich subset of the barley genome confirm that the deconvolution is accurate (almost 70% of left/right pairs in paired-end reads are assigned to the same BAC, despite being processed independently) and the BAC assemblies have good quality (the average sum of all assembled contigs is about 88% of the estimated BAC length). Data availability: Barley raw sequencing data for one set of 2,197 MTP gene-enriched BACs can be obtained from NCBI Sequence Read Archive (http://www.ncbi.nlm.nih.gov/sra?term=(SRA047913))

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Combinatorial Pooling Enables Selective Sequencing of the Barley Gene Space

For the vast majority of species - including many economically or ecologically important organisms, progress in biological research is hampered due to the lack of a reference genome sequence. Despite recent advances in sequencing technologies, several factors still limit the availability of such a critical resource. At the same time, many research groups and international consortia have already...

متن کامل

Deconvoluting the BAC-gene relationships using a physical map.

MOTIVATION The deconvolution of the relationships between BAC clones and genes is a crucial step in the selective sequencing of the regions of interest in a genome. It usually requires combinatorial pooling of unique probes obtained from the genes (unigenes), and the screening of the BAC library using the pools in a hybridization experiment. Since several probes can hybridize to the same BAC, i...

متن کامل

Accurate Decoding of Pooled Sequenced Data Using Compressed Sensing

In order to overcome the limitations imposed by DNA barcoding when multiplexing a large number of samples in the current generation of high-throughput sequencing instruments, we have recently proposed a new protocol that leverages advances in combinatorial pooling design (group testing) [9]. We have also demonstrated how this new protocol would enable de novo selective sequencing and assembly o...

متن کامل

Deconvoluting BAC-Gene Relationships Using a Physical Map

Deconvolution of relationships between bacterial artificial chromosome (BAC) clones and genes is a crucial step in the selective sequencing of regions of interest in a genome. It often includes combinatorial pooling of unique probes obtained from the genes (unigenes), and screening of the BAC library using the pools in a hybridization experiment. Since several probes can hybridize to the same B...

متن کامل

Scrible: Ultra-Accurate Error-Correction of Pooled Sequenced Reads

We recently proposed a novel clone-by-clone protocol for de novo genome sequencing that leverages combinatorial pooling design to overcome the limitations of DNA barcoding when multiplexing a large number of samples on second-generation sequencing instruments. Here we address the problem of correcting the short reads obtained from our sequencing protocol. We introduce a novel algorithm called S...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:
  • CoRR

دوره abs/1112.4438  شماره 

صفحات  -

تاریخ انتشار 2011